A Parallel Learning Algorithm for Text Categorization on PIRUN Beowulf Cluster

نویسندگان

  • Canasai Kruengkrai
  • Chuleerat Jaruskulchai
چکیده

Text categorization is the process of classifying documents into predefined categories or classes based on their content. Since text data rapidly increase on the Internet, the scalability of the algorithm is required to handle such massive data. In this paper, we propose a parallel learning algorithm for text categorization based on the combination of the Expectation-Maximization (EM) algorithm and the naive Bayes. Our experiment performed on a 72 nodes Beowulf cluster called PIRUN. The preliminary experimental results show that our parallel implementation has reasonable speedup characteristics.

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Improving the Operation of Text Categorization Systems with Selecting Proper Features Based on PSO-LA

With the explosive growth in amount of information, it is highly required to utilize tools and methods in order to search, filter and manage resources. One of the major problems in text classification relates to the high dimensional feature spaces. Therefore, the main goal of text classification is to reduce the dimensionality of features space. There are many feature selection methods. However...

متن کامل

Building a Large Scalable Internet Superserver for Academic Services with Linux Cluster Technology

With the speed and bandwidth offered by the next generation Internet technology, there is a need for large and scalable Internet server that can provides an adequate computing power and storage for the new generation Internet applications. This requires a huge investment in a very large and expensive commercial server system. Recently, the emergence of Linux PC clustering or so-called Beowulf C...

متن کامل

Parallel Nearest Neighbour Algorithms for Text Categorization

In this paper we describe the parallelization of two nearest neighbour classification algorithms. Nearest neighbour methods are well-known machine learning techniques. They have been successfully applied to Text Categorization task. Based on standard parallel techniques we propose two versions of each algorithm on message passing architectures. We also include experimental results on a cluster ...

متن کامل

Parallelization of Noise Reduction Algorithm for Seismic Data on a Beowulf Cluster

This paper presents the parallelization of a sequential noise reduction algorithm for seismic data processing into a parallel algorithm. The parallel algorithm was developed using C language with the utilization of the Message Passing Interface (MPI) library. The proposed algorithm has been implemented on an experimental Beowulf cluster which consists of 12 nodes operating on Linux Ubuntu platf...

متن کامل

Solving Traveling Salesman Problem on Cluster Compute Nodes

In this paper, we present a parallel implementation of a solution for the Traveling Salesman Problem (TSP). TSP is the problem of finding the shortest path from point A to point B, given a set of points and passing through each point exactly once. Initially a sequential algorithm is fabricated from scratch and written in C language. The sequential algorithm is then converted into a parallel alg...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2005